AI Learning Series · Part 20

Small Models & On-Device AI

Doc 08 told the GPU-datacenter story; doc 10 the serving stack. This is the mirror image: what happens when the "GPU factory" shrinks into the phone, the laptop, the NPU — same bandwidth wall, a budget 30× smaller.

Long Context & Memory
→
Small & On-Device
→
MoE Frontier

01 The Big Picture

In doc 08 the enemy was bandwidth at 3,000 GB/s of HBM. On your phone the same wall exists — at ~100 GB/s of LPDDR. The physics is identical; the budget is not.

A datacenter H100 reads model weights at ~3,350 GB/s and pays for power after the fact, from a wall socket. A phone reads them at roughly 50–100 GB/s from LPDDR5, and pays during the fact, from a battery, in joules the user can feel as heat. That is the entire design space in one sentence: everything on-device is the datacenter problem with 30× less bandwidth and a power envelope of watts, not kilowatts.

This is why the on-device stack is not "a smaller cloud stack." It is an inverted set of priorities: the cloud optimizes throughput per GPU-dollar; the device optimizes joules per token and latency per tap. Models that look trivially small by datacenter standards (7B parameters) are, on a phone, a 3.5 GB weight file that must stream through a narrow memory pipe for every single token.

🔄
The flip: doc 08 said "decode is memory-bound — FLOPs are abundant, bytes are scarce." On-device that statement becomes brutal. An NPU may advertise 60 TOPS of compute, but LPDDR feeds it ~100 GB/s — so the silicon idles, waiting on bytes, exactly like its datacenter sibling. Only the numbers shrank; the law didn't.

02 What It Is — The Edge Inference Envelope

On-device inference is running the full forward pass — weights, KV cache, sampling — on the user's own silicon: phone NPU, laptop CPU/GPU, or a browser. The defining discipline is the envelope: the fixed set of resources the model must live inside.

Memory envelope

iOS gives an app a few GB; Android often less. Weights + KV cache + activations + the rest of your app must all fit. 4-bit 7B ≈ 3.5 GB of weights alone — that is the model budget for most phones.

Bandwidth envelope

LPDDR ~50–100 GB/s vs HBM ~3,000 GB/s. Since decode speed is bytes-per-second over bytes-per-token, this single number sets your tokens/s ceiling before any software runs.

Power envelope

Sustained NPU budgets are ~1–5 W. Sustained load throttles; the phone gets warm and the OS will slow you down. You design for joules per inference, not tokens per second per dollar.

The crucial reframing for engineers: the envelope is bytes, not FLOPs. Every architectural choice on-device — quantization, GQA, small context windows, weight caching in SRAM — is a strategy for reading fewer bytes per token. Compute is nearly free at these scales; memory traffic is the whole bill.

03 Why It Exists

Given how good cloud APIs are, why compress models into a phone at all? Four pressures, none of which the cloud can ever fully answer:

Privacy is physics, not policy. Data that never leaves the device cannot be logged, subpoenaed, or breached in transit. A medical draft, a keyboard's autocomplete, a camera's scene text — the only zero-trust answer is zero bytes on the wire.
Latency and always-on. No round trip: first token in tens of milliseconds, not 300+. That's the difference between autocomplete that feels psychic and one that feels laggy. And inference that runs with the screen off — categorization, triage, wake-word-adjacent tasks — must be free of network and cheap enough to run constantly.
Offline is a feature. Airplane mode, rural connectivity, enterprise air-gapped sites. A feature that dies without signal is a support ticket; one that degrades gracefully is a selling point.
Cost flips sign. In the cloud every token is a bill (doc 07's currency). On-device the marginal cost of a token is a fraction of a joule — effectively free. That economics unlocks products where the model runs millions of tiny times per day, which no per-token pricing could survive.

04 How It Works — One Token, From Battery to Byte

Follow a single decode step through the device's memory hierarchy — then watch what "send to cloud instead" would have cost.

user taps prompt tokenized on-device → token "…" needs next-token decode LPDDR — 3.5 GB of 4-bit weights ~100 GB/s · every layer read EVERY token this stream is the whole cost SRAM (on-chip) — hot slab current layer weights + KV window ~100× cheaper per byte than DRAM reads NPU: multiply-accumulate 60 TOPS on paper — mostly idle, waiting for the LPDDR stream power bill = bytes × energy/byte DRAM read ≈ 10 pJ/B · SRAM ≈ 0.1 ~1–3 J per second of decode decode speed ceiling 100 GB/s ÷ 3.5 GB/token ≈ 28 tokens/s — the upper bound THE ALTERNATIVE — same prompt, sent to cloud: LOCAL: bytes stay on device · private · ~28 tok/s · ~free drafts, autocomplete, classification, rewriting CLOUD: prompt leaves the device · 300+ ms RTT · per-token bill frontier quality, 100B+ params, instant scale-up A privacy gesture — the lock icon in your keyboard — is a routing decision between these two rows. The best products do both: draft locally, verify in the cloud (section 07).

Read the loop again as a system: every token forces a full sweep of the weights — LPDDR → SRAM → NPU → logits — then the whole cycle repeats. The NPU is not the bottleneck; the pipe is. Which is why the next two sections are almost entirely about shrinking the pipe: fewer bytes per parameter (quantization), fewer parameters per token (distillation, sparsity), fewer bytes per attention step (GQA, small context — doc 22).

05 The Memory Math

Three formulas decide the entire on-device design space. All three are byte-counting; none involve FLOPs.

1 · Weight footprint

bytes(weights) = N_bits · P / 8 7B params @ 4-bit → 4 × 7×10⁹ / 8 = 3.5 GB 7B params @ 8-bit → 7.0 GB (often exceeds the app budget)

This is why quantization is not an optimization on-device — it is admission criteria. At 16-bit the same model is 14 GB and simply cannot ship.

2 · KV cache footprint

bytes(KV) = 2 · s · L · n_kv · d · bytes_per_elem s = 4096 ctx · L = 32 layers · n_kv = 4 (GQA) · d = 128 · fp16 → 2 × 4096 × 32 × 4 × 128 × 2 = 268 MB (0.25 GB)

That "2" is K and V; n_kv = 4 instead of 32 is Grouped-Query Attention — 8× less cache than classic MHA. Without GQA the same context costs 2 GB: a second model's worth of RAM. Phones need GQA + short contexts for the same reason datacenters do (docs 08, 22) — the KV cache is read on every decode step too.

3 · Decode speed

tokens/s ≈ BW_effective / bytes_read_per_token ≈ BW_effective / (N_bits · P / 8) LPDDR 100 GB/s, 4-bit 7B (3.5 GB/token): → 100 / 3.5 ≈ 28 tok/s ← the hard ceiling HBM 3,350 GB/s, same model: → ~957 tok/s

Instant corollary: dropping weights from 8-bit to 4-bit halves bytes-per-token and doubles decode speed — quantization is a latency feature, not just a memory one, because decode is memory-bound (doc 08's law, restated in doc 10).

QuantityFormula7B @ 4-bit on phoneSame model, datacenter
WeightsN_bits·P/83.5 GB3.5 GB (of 80 GB HBM)
KV cache @ 4k ctx2·s·L·n_kv·d·20.25 GB (GQA)0.25 GB (2 GB w/o GQA)
Memory BW—~100 GB/s~3,350 GB/s
Decode ceilingBW / bytes-per-token~28 tok/s~957 tok/s
Power sourcebytes × pJ/bytebattery, ~1–3 W sustainedwall socket, ~700 W
📐
Sanity check every on-device pitch with row 4: if a demo claims 100 tok/s from a 7B model at 4-bit on LPDDR, the weights can't be read from DRAM every token — something else is happening (smaller model, speculative decoding, or caching tricks). The bandwidth formula doesn't negotiate.

06 Quantization & Distillation — The Tooling That Shrinks P

Sections 05's formulas have two levers: N_bits (quantization) and P (distillation/pruning). These are the toolchains that make a phone-sized envelope viable at all.

Per-channel quantization

symmetric: W ≈ s · Q, s = max|W_row| / (2ᵇ⁻¹ − 1) asymmetric: W ≈ s · Q + z (zero-point z for skewed distributions) group size 32–64: one scale (s, z) per 32–64 weights, not per row

Storing one scale per output channel works for activations and matrices, but LLM weight distributions have outlier channels that poison whole rows. Group-wise scales (every 32 or 64 weights) localize the damage. NF4 (NormalFloat4) goes further: instead of uniformly spaced 4-bit levels, it uses quantile-spaced levels of a normal distribution — the values real weights actually take — so each code is equally likely and information-per-bit is maximal.

GPTQ / AWQ — error-aware rounding

Naive rounding quantizes each weight independently. GPTQ and AWQ instead ask: what matters is the layer's output. The objective is to minimize the output error, weighted by the activation statistics of real data:

minimize ‖ W·X − Q(W)·X ‖² (X = calibration activations)

GPTQ compensates rounding errors of one weight by adjusting its not-yet-quantized neighbors, column by column. AWQ observes that a few salient channels dominate the loss, protects their precision, and pushes more aggressive bits onto the rest. Result: 4-bit models within a point or two of fp16 on benchmarks — the difference between "the math says 3.5 GB" and "the model still works at 3.5 GB."

Structured sparsity — 2:4 on tensor cores

every contiguous group of 4 weights → force 2 to zero → store 2 values + 2-bit indices storage: 2 × 8 bits + 2 × 2 bits = 20 bits / 4 weights = 5 bits/weight (vs 8) compute: NVIDIA sparse tensor cores skip zeros → 2× MAC throughput, free

Sparsity is quantization's sibling: another way to trade a little quality for fewer bytes and fewer operations — and unlike unstructured pruning, the "2 of 4" pattern is regular enough for hardware to exploit without gather/scatter logic.

Distillation — shrinking P itself

L_student = KL( softmax(teacher_logits / T) ‖ softmax(student_logits / T) )

The student learns from the teacher's softened distribution. With temperature T > 1, probabilities sharpen less: instead of "next token: 99.9% 'the'," the teacher reveals "…and 0.02% 'a', 0.01% 'my'" — the dark knowledge of how tokens relate. That inter-class structure is worth thousands of hard-label examples, and it's why a 1–3B student distilled from a 70B teacher punches far above its parameter count. Distillation and quantization compose: distill to 3B, quantize to 4-bit → 0.75 GB, comfortably on-phone.

✓ Do

Quantize with calibration data from your domain; keep embeddings/lm_head at higher precision; use group size 32–64 for 4-bit; distill before quantizing; measure quality on your tasks, not MMLU alone.

✗ Don't

Go below 3–4 bits on small models without a proven recipe (quality collapses non-linearly); quantize KV cache below 8-bit blindly (doc 22); assume fp16 benchmarks predict INT4 behavior; prune unstructured and hope hardware notices.

07 Hybrid Architectures — Edge Draft, Cloud Verify

Pure-local and pure-cloud are the endpoints of a spectrum; shipping products live in the middle. Two patterns dominate.

Speculative edge→cloud

The on-device small model drafts the next several tokens instantly (free, private, offline); the cloud model verifies them in one parallel prefill pass — exactly the draft/verify split of doc 24's speculative decoding, stretched across the network boundary. Acceptance costs the cloud ~1 forward pass for k drafted tokens; rejection falls back to cloud-native decoding. The device pays zero cloud latency for accepted spans, and drafts never containing sensitive text needn't be sent at all.

Router: who answers this prompt?

A small π_router — often the on-device model itself, or a classifier — grades each request by expected difficulty and routes:

U(route) = q(route) · v(task) − c(route) q = expected quality · v = value of being right · c = cost (latency + ¢ + joules + privacy) route to local if U(local) ≥ U(api) — i.e., easy tasks with privacy value

Simple queries ("is this email rude?") stay local: q_local ≈ q_api, c_local ≈ 0. Frontier reasoning goes to the API. The router is a mixture-of-tokens at the product scale — the same "spend cheap tokens on easy work" logic as MoE's expert routing, applied to the edge–cloud boundary. (MoE itself is the next doc's subject.)

🧭
Where routing already lives: mobile keyboards (local next-word, cloud rewrite), photo search (local embeddings, cloud generation), and every "on-device AI" badge are π_router in production. The lock icon users trust is literally this routing decision rendered as a gesture.

08 Engineering Takeaways — Joules and Thermals

On the device, the unit of cost is the joule. Power per inference decomposes the same way speed did — into memory traffic:

E(inference) ≈ bytes_moved × energy_per_byte + E_compute DRAM read ≈ 10 pJ/byte · SRAM access ≈ 0.1 pJ/byte (≈100×)

That 100× gap is why NPU architects obsess over dataflow: keep the current layer's weights and KV window resident in SRAM and reuse them across timesteps; touch LPDDR as few times as possible. A model that fits entirely in SRAM runs orders of magnitude cheaper per token than one streaming from DRAM — but SRAM is measured in MB, not GB, so most LLMs live in the DRAM-streaming regime and must "get few external reads" via aggressive batching of reuse. Apple and other NPUs pair ~60 TOPS of compute with exactly this SRAM-first design — yet the LPDDR ceiling of ~100 GB/s still caps end-to-end decode, per section 05.

Budget in J/inference. A background task at 2 J/run × 1,000 runs/day = 2,000 J ≈ 2% of a phone battery — fine. The same task every keystroke is not. Measure per-workload, not per-second.
Design for throttle, not peak. Sustained decode heats the die; the OS drops clocks within seconds. Benchmark "tokens 30–120," not the first burst. Short generations that finish before the throttle curve bites feel faster than long ones that crawl.
Prefer fewer bytes over more cores. Because both speed and energy are byte-proportional, quantization and context-trimming improve battery and latency simultaneously — the rare free lunch in this doc.

09 Mental Models

On-device LLM = disk-backed database; cloud = warehouse VM

Local inference is SQLite: small, private, instant, always there, bounded by your hardware. The API is a managed warehouse: vast, elastic, metered per query, and reachable only over the network. Products, like apps, usually need both. Lets you reason about: why "local vs cloud" is a storage-engine choice, not a religion.

A database serves arbitrary queries from an index; a model serves from fixed weights — the local one can't "look up" what it was never trained on.
The straw vs the firehose

HBM is a firehose feeding the datacenter GPU; LPDDR is a straw feeding the NPU. The model is the same liquid — the only question is how fast you can drink. Quantization is thinning the liquid (fewer bits per drop); GQA is a smaller mouthful per sip. Lets you reason about: why every on-device technique is "drink less," never "sip harder."

Thinning too far changes the taste — below ~3–4 bits, quality doesn't degrade, it collapses.
SRAM is your desk, LPDDR is the library

Anything on your desk (SRAM) is instant to consult but the desk is tiny; the library (LPDDR) is huge but every trip costs a walk. Good NPU code is a good researcher: bring a chapter to the desk, work through it thoroughly, minimize library trips. Lets you reason about: the 100× energy gap and why dataflow matters more than TOPS.

Unlike a researcher, decode must re-read every shelf (all weights) for every single token.

10 Common Misconceptions

"More TOPS = faster LLM." Decode is memory-bound; a 60-TOPS NPU streaming 3.5 GB through a 100 GB/s straw still lands at ~28 tok/s. TOPS matter for prefill and CNNs, not token generation.

"Quantization just loses a little accuracy." Done right (GPTQ/AWQ, group-wise, calibrated), 4-bit costs a point or two. Done naively (per-tensor absmax, no calibration), it can destroy a model. The variance is in the method, not the bit count.

"On-device means the model is 'as smart as' a small cloud model." Parameter count transfers only with equal context, sampling, and tooling. On-device models run shorter contexts (KV math), smaller output budgets, and no retrieval — judge them on the envelope they actually run in.

"Everything should move on-device for privacy." Routing is per-task. A request with zero sensitive content gains nothing from local execution but pays in quality. The right primitive is the router with a privacy term in its utility function, not blanket local-first.

"Browsers can't run real models." WebGPU gives the tab near-native access to the GPU; WASM+SIMD covers CPU fallback. Multi-megabyte GGUF-style weight formats stream and run in-tab — the browser is now a legitimate inference target, with zero install and the sandbox as the privacy boundary.

🗺️
Where this leaves us: the device gives you privacy, latency, and free tokens — in exchange for a 3.5 GB, ~28 tok/s, joule-metered envelope, managed with quantization, distillation, GQA, and routing. Next: what if instead of routing between models, one model routed inside itself — MoE, the frontier of sparse compute. Start from doc 08 to see the same bandwidth wall at datacenter scale, and doc 19 for how long contexts strain this exact envelope.